Papers by Ihsan Ayyub Qazi

4 papers
Language Model-Driven Data Pruning Enables Efficient Active Learning (2026.findings-eacl)

Copied to clipboard

Challenge: Existing data pruning methods for active learning are expensive and time-consuming.
Approach: They propose a plug-and-play data pruning strategy that leverages language models to prune the unlabeled pool.
Outcome: The proposed pruning strategy outperforms existing pruning methods on translation, sentiment analysis, topic classification, and summarization tasks on diverse datasets.
POLAR: A Benchmark for Multilingual, Multicultural, and Multi-Event Online Polarization (2026.findings-acl)

Copied to clipboard

Challenge: polarization is a pervasive threat to democratic institutions, civil discourse, and social cohesion worldwide . most existing datasets focus on English or high-resource languages, reflecting a widespread trend across NLP tasks .
Approach: They propose a multilingual, multicultural, and multi-event dataset with over 110K instances in 22 languages drawn from diverse online platforms and real-world events.
Outcome: The proposed dataset analyzes polarization detection, type, and manifestation using a variety of annotation platforms adapted to each cultural context.
Deepfake Defense: Constructing and Evaluating a Specialized Urdu Deepfake Audio Dataset (2024.findings-acl)

Copied to clipboard

Challenge: Automatic speaker verification systems are facing escalating challenges due to deepfake attacks.
Approach: They propose a Urdu deepfake audio dataset for deepfak detection focusing on two spoofing attacks – Tacotron and VITS TTS.
Outcome: The proposed dataset evaluates two spoofing attacks in Urdu with a human evaluation to gauge whether people are able to distinguish deepfake audios from real (bonafide) audios.
To Label or Not to Label: Hybrid Active Learning for Neural Machine Translation (2025.coling-main)

Copied to clipboard

Challenge: Active learning (AL) techniques reduce labeling costs for training neural machine translation models by selecting smaller representative subsets from unlabeled data for annotation.
Approach: They propose an AL strategy that combines uncertainty and diversity for sentence selection.
Outcome: The proposed method prioritizes diverse instances having high model uncertainty for annotation in early iterations.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations